Back

Journal of Genetics and Genomics

Elsevier BV

Preprints posted in the last 7 days, ranked by how well they match Journal of Genetics and Genomics's content profile, based on 38 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
Eucalyptus microRNA Archive (EMA): a multi-study and cross-condition curated database of microRNAs in Eucalyptus grandis

Aires Teixeira, J. V.; Motta Venancio, T.; Quintanilha-Peixoto, G.; Pimenta de Oliveira, K. K.

2026-08-31 plant biology 10.64898/2026.08.29.747619 medRxiv
Top 1%
0.5%
Show abstract

MicroRNAs (miRNAs) are key post-transcriptional regulators of development, stress response, and secondary cell wall formation in woody plants, yet annotations for Eucalyptus grandis, the world's most widely planted hardwood, remain fragmented across studies using incompatible discovery pipelines and filtering criteria. Here we present the Eucalyptus MicroRNA Archive (EMA), a curated, locus-resolved database integrating three independent small RNA sequencing datasets spanning vegetative tissue, somatic embryogenesis, and mechanically induced tension wood formation. Applying annotation criteria aligned with current plant miRNA standards, EMA catalogs 99 curated miRNAs (31 previously described, 68 novel) organized into 34 family-level groupings under a three-tier confidence system, known-reference-supported, multi-study replicated, or single-study, that preserves study-of-origin and sample-level evidence for every entry. Cross-study comparison showed that only 9 of 99 entries (9.1%) were independently supported by all three datasets, supporting an evidence-tiered rather than binary annotation scheme. Target prediction against the E. grandis transcriptome yielded 1,773 miRNA-target interactions spanning 764 loci, integrated into a combined miRNA-target and protein-protein interaction network. This network resolved into functionally coherent, mutually isolated clusters, including an miR482-associated NBS-LRR/TIR disease-resistance hub with a substantial translational-repression component, alongside modules enriched for ribosome biogenesis and translation, DNA replication, and nitrogen and carbohydrate metabolism. EMA is publicly accessible through an interactive web dashboard, with all curated data, source code, and analysis scripts openly available, providing a reproducible, extensible framework for E. grandis miRNA research and a template for similarly structured resources in other non-model woody species.

2
A Curated Pharmacogenomic Allele Catalog for Sub-Saharan African Populations

SULAIMAN, M. A.; Oyeyemi, B. F.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.25.26361354 medRxiv
Top 1%
0.4%
Show abstract

Sub-Saharan African populations carry pharmacogenomic alleles poorly represented in the European-derived reference panels underlying most clinical genotyping tools. We present a curated, machine-readable catalog of nine actionable alleles across six pharmacogenes (CYP2D6, CYP2B6, CYP2C9, CYP2C19, CYP3A5, NAT2) with African-specific frequency ranges, functional annotations, and evidence levels derived from reanalysis of 661 high-coverage whole-genome sequences across seven 1000 Genomes Project African populations. Direct comparison against PharmCAT v3.4.0 shows that CYP2D6 produces zero diplotype calls (0/661 samples callable) due to monomorphic reference positions absent from standard variant-only VCF output, a known limitation whose consequences for African allele carriers had not been reported. afripharmagen's reduced-position strategy identifies 243 CYP2D617 and 134 CYP2D629 carriers from the same input. For CYP2B6, CYP2C9, CYP2C19, and NAT2, both tools show concordance of 95-100%. Frequency gradients (CYP2B66: 30-50%; CYP2D617: 15-35% in West Africa; CYP3A5*1: 60-95%) translate directly into prescribing risk for efavirenz, tramadol, tacrolimus, and isoniazid. Pharmacogenomic decision support in African settings must incorporate population-specific allele definitions and input-format-aware strategies.

3
A germline KDM3C polymorphism impairs DNA repair and sensitizes to chemoradiotherapy

Hasan, A.; Demidova, E. V.; Priyadarshini, P.; Czyzewicz, P.; Gathuka, L.; Murayama, T.; Zhou, Y.; Kiss, Z. A.; Shastry, R. K.; Andrake, M.; Hearne, G.; Devarajan, K.; Wu, C.; Shah, A.; Schultz, B. M.; Connolly, D. C.; Rosen, G. L.; Canadas, I.; Liu, J. C.; Burtness, B. A.; Smith, J. J.; Dunbrack, R. L.; Golemis, E. A.; Whetstine, J. R.; Meyer, J. E.; Arora, S.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.26.26360896 medRxiv
Top 2%
0.3%
Show abstract

Chemoradiotherapy (CRT) is the standard-of-care therapy for many solid malignancies, yet predictive biomarkers of treatment response remain limited. We identified a germline single nucleotide polymorphism (SNP) in an intrinsically disordered region of the lysine demethylase KDM3C/JMJD1C (p.S464T) that is associated with CRT outcomes in locally advanced rectal cancers (LARC) and head and neck squamous cell carcinoma (LA-HNSCC). In silico modeling with AlphaFold predicted S464T substitution influenced interaction between phosphorylated KDM3C and RNF8 FHA domain. In cellular models, conversion of S464 to T464 increased sensitivity to DNA-damaging agents. S464T substitution impaired damage-induced MDC1-RAP80 signaling and downstream RAP80-BRCA1 colocalization. SNP carrying cells impaired DNA repair causing genotoxic stress that is associated with increased cGAS-cGAMP innate immune signaling and increased apoptosis. Population analyses with the SNP highlighted an increase incidence of UV-induced skin and other cancers, linking inherited variation in the chromatin regulatory gene KDM3C to genome instability, cancer risk, and therapeutic vulnerability.

4
A 515,579-Genome Reference Panel Improves Rare-Variant Imputation Across Multiple Underrepresented Populations

Ivankovic, F.; Ko, A.; Aster, M. M.; Balaconis, M. K.; Banks, E.; Bemis, M.; Cibulskis, K. R.; Degatano, K.; Gauthier, L. D.; Grant, G.; Hatcher, A.; Kachulis, C.; Karczewski, K. J.; Labrecque, S. M.; Lawson, J.; Liao, C.; Magner, R.; Munshi, R.; Schatz, M. C.; Schultz, P. M.; Shah, S. P.; Sheets, E. A.; Tibbetts, K.; Vernest, K. A.; Ye, R.; Gabriel, S.; Lennon, N. J.; Neale, B. M.; Browning, B. L.; Lichtenstein, L. T.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.25.26361247 medRxiv
Top 2%
0.3%
Show abstract

Genotype imputation remains essential for large-scale human genetics studies, but its performance is limited by the size and ancestral diversity of available reference panels, reducing accuracy for rare variants and underrepresented populations. Here, we present a cloud-based imputation service built on a multi-ancestry reference panel derived from 515,579 jointly phased genomes from the All of Us (N=414,830) and National Human Genome Research Institute's Analysis, Visualization, and Informatics Lab-space (AnVIL, N=100,749) datasets. The All of Us + AnVIL reference panel is highly diverse and includes 261,163 participants most genetically similar to non-European reference populations, spanning 665,398,839 high-quality autosomal sites, representing a nearly 50% increase over TOPMed, the previous largest imputation service. Across multiple ancestry groups, the panel enables accurate imputation (empirical R2 0.8) for variants with allele frequencies as low as 0.2%, extending reliable imputation into the rare-variant frequency spectrum, including allele frequencies down to 0.002% and 0.006% for samples with European ancestry and African ancestry in the United States, respectively. Compared with TOPMed, the panel improves imputation accuracy across all ancestry groups except Africans, and recovers additional trait-associated variants not represented in existing reference panels. To facilitate broad community access while preserving participant privacy, we deploy the panel through a secure cloud-based imputation platform using privacy-preserving recombined haplotypes. This resource establishes a new foundation for genome-wide association studies (GWAS) and fine-mapping, especially in previously underrepresented populations.

5
A subgenome-resolved and chromosome-scale reference genome assembly of allotetraploid wheat wild relative Aegilops peregrina

Singh, J.; Gudi, S.; Maughan, P. J.; Gill, U.; Gupta, R.

2026-08-30 genomics 10.64898/2026.08.28.747929 medRxiv
Top 2%
0.3%
Show abstract

Aegilops peregrina is a wild allotetraploid wheat wild relative and an important source of genetic diversity for stress tolerance and agronomic traits. Here, we report a subgenome-resolved, chromosome-scale reference genome assembly of a drought tolerant and stem rust resistant Ae. peregrina accession PI 604178 generated using PacBio HiFi and Hi-C sequencing. The 10.13 Gb assembly contains 98.81% of sequence anchored to 14 pseudomolecules representing the seven S and seven U chromosomes, with contig and scaffold N50 values of 25.84 and 746.48 Mb, respectively. The assembly achieved a consensus quality value of 74.61, 97.83% k-mers completeness, and 99.9% BUSCO completeness. LTR Assembly Index values of 20.43 and 18.79 for the S and U subgenomes, respectively, further supported high continuity across repeat-rich regions. Repetitive elements comprise 85.93% of chromosome-anchored assembly. We annotated 59,910 high-confidence protein-coding genes, with comparable gene representation across the two subgenomes. This reference genome provides a high-quality genomic framework for comparative analyses, characterization of important loci regulating agronomic and resilience related traits, and sequence-guided exploitation of Ae. peregrina allelic diversity for wheat improvement.

6
BRIX1 Promotes Hepatocellular Carcinoma Progression via the MAPK/ERK Pathway and Serves as a Prognostic Biomarker

Pan, X.; Wang, x.; Zhou, Y.

2026-08-31 cancer biology 10.64898/2026.08.26.747409 medRxiv
Top 3%
0.2%
Show abstract

Hepatocellular carcinoma (HCC) is particularly aggressive and difficult to treat. Due to the lack of early clinical diagnosis and the unsatisfactory clinical treatment effect, it is particularly important to identify novel markers that can predict tumor behavior in HCC. biogenesis of ribosomes BRX1 (BRIX1) is abundant in various tissues of the human body. However, the regulatory mechanisms and its role in various tissues are not fully understood. Here, we analyzed the expression pattern of BRIX1 in HCC from public gene expression databases and tissue samples from clinical HCC. We confirmed that BRIX1 was upregulated in both HCC cell lines and HCC paraffin section samples. BRIX1 depletion significantly dicreased the capacity of cells to grow and migrate in vitro, and knockdown BRIX1 suppressed tumor growth in xenograft tumor model. Mechanistically, BRIX1 depletion suppressed the MAPK/ERK pathway, as reflected by reduced phosphorylated ERK (p-ERK) levels. In summary, we provide a rational clue for the further investigation of BRIX1 as an invaluable biological marker for diagnosing and predicting prognosis of patients with HCC.

7
The first chromosome-scale genome assembly of Blumeria graminis f. sp. avenae provides insights into genome evolution and host specialization

Ding, Y.; Zhang, P.; Ociepa, T.; Nucia, A.; Guan, H.; Kowalczyk, K.; Park, R. F.; Okon, S.

2026-08-30 genomics 10.64898/2026.08.28.747853 medRxiv
Top 3%
0.2%
Show abstract

Blumeria graminis f. sp. avenae (Bga), the causal agent of oat powdery mildew, is one of the most host-specialized members of the B. graminis species complex. Despite its agricultural importance, the lack of a high-quality reference genome has limited studies of host specialization, virulence evolution and comparative genomics in this pathogen. Here, we generated the first chromosome-scale genome assembly of Bga using an integrative approach combining long- and short-read sequencing, Hi-C scaffolding and transcriptome data. The Bga genome exhibits hallmark features of powdery mildew fungi, including extensive repeat content and low gene density. Comparative analyses revealed that genome expansion is primarily associated with historical transposable element proliferation rather than recent transpositional activity. Genome organization is consistent with a functionally stratified "one-speed" model, in which genes associated with pathogenicity, including predicted effectors and infection-responsive genes, are preferentially located in transposable element-rich regions characterized by reduced synteny conservation and extended intergenic spaces. In contrast, conserved genes are concentrated in compact genomic regions and maintain strong syntenic conservation across cereal-infecting formae speciales. Hi-C analyses demonstrated a highly structured chromatin architecture and revealed genome organization patterns associated with infection-related gene expression. Comparative genomic analyses indicated that host specialization in Bga is driven by localized diversification of a relatively small subset of genes rather than large-scale genome restructuring. These results provide the first high-quality genomic resource for Bga and offer new insights into the evolutionary mechanisms underlying host specialization in powdery mildew fungi.

8
MetaDome 2027: a comprehensively updated resource for aggregating missense variant evidence across homologous human protein domains

Wiel, L.; Ferraro, F.; Yu, J.; Zhen, J.; Nachun, D.; Mendez, R.; Reuter, C. M.; Cui, J. L.; Bonner, D. E.; Carter, J. N.; Marwaha, S.; van de Vorst, M.; Emami, S.; Kravets, E.; Neu, M. B.; van Ham, T. W.; Kleefstra, T.; Ashley, E. A.; Bernstein, J. A.; Montgomery, S. B.; Gilissen, C.; Wheeler, M. T.

2026-08-31 bioinformatics 10.64898/2026.08.26.747388 medRxiv
Top 3%
0.2%
Show abstract

The interpretation of missense variants remains a major challenge in clinical genetics. "Meta-domains" aggregate population and pathogenic variation across homologous Pfam domain instances in the human proteome, providing per-residue context for interpreting variants of uncertain significance (VUS). Our 2019 implementation, MetaDome, is widely used and named in clinical variant-classification guidelines. Here we present the MetaDome 2027 update, featuring a comprehensively updated dataset and GRCh38 support. The redesigned pipeline enables incremental updates of GENCODE, UniProtKB/Swiss-Prot, Pfam, gnomAD, and ClinVar while maintaining 100% sequence-identity gene-to-protein mapping. Annotated Pfam domain instances grew 14.9% from 71,419 to 82,069 and meta-domain-eligible Pfam families ([≥]2 human occurrences) by 73.3% from 3,334 to 5,778; Pfam domains are annotated to 92% of human proteins. Approximately 43% of mapped protein-coding nucleotides (14.3 million in GRCh38, 13.8 million in GRCh37) are in a meta-domain; in GRCh38 67.9% (37,692 of 55,548) of pathogenic or likely pathogenic ClinVar missense variants fall at such a position. We show how MetaDome helped reclassify a de novo missense VUS in RALA and identify 52,463 ClinVar missense VUS for which meta-domains supply otherwise unavailable pathogenic evidence. MetaDome is freely available at www.metadome.app.

9
Generation and characterization of a patient-specific human induced pluripotent stem cell line from a Skogholt syndrome patient (ASCFi003-A)

Przybyla, W.; Gupta, S.; Fjerdingstad, H. B.; Selnes, P.; Sharma, K.

2026-08-31 cell biology 10.64898/2026.08.29.747981 medRxiv
Top 4%
0.1%
Show abstract

We report the generation and characterization of a human induced pluripotent stem cell (iPSC) line derived from dermal fibroblasts of a patient with Skogholt disease, a rare maternally inherited neurodegenerative syndrome associated with choroid plexus dysfunction and impaired cerebrospinal fluid (CSF) homeostasis. Patient fibroblasts were reprogrammed using the non-integrating Repro-OSKGM kit. The resulting iPSC line exhibited typical pluripotent morphology, expressed canonical pluripotency markers, maintained a normal karyotype, retained the disease-associated genetic variant, was mycoplasma-free, and demonstrated trilineage differentiation potential. We also made choroid plexus (ChP) like organoids from the generated iPSCs. This patient-specific iPSC line provides a valuable resource for generating choroid plexus organoids and neurons to investigate disease mechanisms and develop therapeutic strategies.

10
Two evolutionary histories in one nucleus: genome remodeling and allelic regulation underlying heterosis in hybrid oil palm

Su, X.; Peng, Y.; Yang, X.; Zhang, F.; Xu, Q.; Ma, Z.; Dong, Y.; Zhou, L.; Xue, H.; Cao, X.; Zou, Z.; Wang, Y.; Zhou, Y.; Zeng, X.

2026-08-31 genomics 10.64898/2026.08.27.747553 medRxiv
Top 5%
0.1%
Show abstract

Oil palm (Elaeis) is the primary source of global vegetable oil. Interspecific hybrids of Elaeis exhibit pronounced heterosis by integrating two distinct subgenomes into a single nucleus, effectively combining the high yield of African oil palm (E. guineensis) with the high unsaturated fatty acid content and disease resistance of American oil palm (E. oleifera). However, the genetic basis underlying heterosis is still unclear. Here, we combine phased genome assembly, comparative genomics, evolutionary genomics and haplotype-aware transcriptomics to unravel the genetic architecture of heterosis of hybrid oil palm. We assemble the highly heterozygous F1 genome ('Reyou 40', 3.75% heterozygosity) into a complete 1.73 Gb T2T haplotype (HapG) and a 1.84 Gb near-T2T haplotype (HapO with17 gaps). Despite 91.56% sequence identity, HapG and HapO diverged in LTR-RT occurrence and PAV affected genes, showing complementary biases in lipid metabolism and stress responses, respectively. Evolutionary genomics revealed that ancient WGDs preserved the palm family. Whereas lineage-specific lipid-related gene expansions in oil palm. Six ancient introgressed regions (~64 Mb) in HapG were reshaped by transposable elements and tandem duplication, showing an enrichment of genes related to resistance and lipid metabolism. Transcriptomically, 82.2% of allelic gene pairs maintained balanced expression, accompanied by parental functional complementarity and dosage buffering, revealing a potential regulatory basis for coordinating parental genetic differences in the hybrid genome. These haplotype-resolved genomic resources offer vital targets for understanding heterosis and accelerating oil palm molecular breeding.

11
Cell-type-resolved somatic variant discovery from bulk long-read sequencing

Fu, Y.; Morley, C.; Masters, L. M.; English, A. C.; Zhu, Y.; Moller, A. G.; Paulin, L. F.; Thompson, B.; Kalef-Ezra, E.; Weissenberger, G.; Shen, H.; Meridith, M.; Manini, A.; Horner, D.; Reed, X.; Muzny, D.; Jaunmuktane, Z.; Khan, Z. M.; Mehta, H.; Timp, W.; Billingsley, K.; Erwin, G. S.; Proukakis, C.; Sedlazeck, F. J.

2026-09-04 genetic and genomic medicine 10.64898/2026.09.01.26361966 medRxiv
Top 5%
0.1%
Show abstract

Somatic mutations arise throughout life, with functional consequences tied to the cell populations in which they occur. Genome-wide studies measure somatic variations in bulk tissue, whereas single-cell approaches resolve cell identity but provide limited sensitivity for complex alleles. Here we developed SniffCell, which uses DNA methylation carried on native long reads to assign somatic variant-supporting molecules to methylation-resolvable cell types. SniffCell builds cell-type-discriminatory methylation signatures across eight tissues, assigns long reads to cell types, and provides cell-type-specific variant calling. Across peripheral blood mononuclear cells and brain benchmarks, SniffCell recovered sorted cell identities and validated cell-type-specific variant assignments using purified immune-cell, neuronal, and oligodendrocyte fractions. In blood, SniffCell recovered lineage-restricted antigen receptor rearrangements and localized a somatic tandem-repeat expansion to T cells. In the frontal cortex, SniffCell identified recurrent neuron-specific tandem-repeat expansions in genes including FGF14, LRRC7 and SH3RF3. Across three brain cohorts comprising 172 donors, recurrent neuron-associated expansions were enriched for GAA-rich motifs. In donors with matched blood, and diverged more strongly from the inherited repeat length, whereas oligodendrocyte-associated alleles more often tracked it. SniffCell transforms native bulk long-read genomes into a cell-type-aware resource for somatic variant discovery and reveals recurrent somatic instability in human tissues at cell-type resolution.

12
In vivo multimodal lineage tracing of mammalian development by DeepTrack barcoding

Guo, C.; Jiang, J.; Wang, X.; Huang, X.; Zhang, S.; Shao, C.; Zhang, M.; Hu, X.; Yang, W.; Shang, F.; Wang, X.; Zhai, H.; Du, Q.; Liu, F.; He, D.; Liu, X.; Peng, G.; Cheng, S.; Zhang, Y.; Pei, D.; Pei, W.

2026-08-31 developmental biology 10.64898/2026.08.29.748052 medRxiv
Top 5%
0.1%
Show abstract

A comprehensive recording of cell fate transitions and underlying molecular changes remains a fundamental goal in developmental biology. Here, we present DeepTrack, a lineage tracing mouse model that integrates in situ cellular barcoding with high-throughput, single-cell multi-omics to simultaneously profile clonal fates, transcriptomic states, and chromatin accessibility. Using DeepTrack, we profiled clonal behaviors during gastrulation and early organogenesis, uncovered early fate priming within epiblast clones, and revealed clonal architecture within distinct regions of the nervous system. Embryo-wide multi-omic lineage tracing at single-cell resolution revealed transcriptional and epigenetic programs underlying fate commitment in neuromesodermal progenitors (NMPs). Clonal tracing with multi-omic profiles enabled inference of fate-associated gene-regulatory networks and identified the transcription factor Cdx2 as a key regulator of mesodermal specification in NMPs. Genetic perturbation of Cdx2 in chimeric embryos impaired paraxial mesoderm differentiation. Together, DeepTrack provides a versatile framework for decoding multimodal regulation of cell fate across diverse developmental contexts.

13
GDNF enemas improve epithelial and immune defects in both aganglionic and ganglionic colon of Hirschsprung mice

Lassoued, N.; Trudel, J.; Lefevre, M.; Gary, A.; Guo, Z.; Yero, A.; Jenabian, M.-A.; Soret, R.; Pilon, N.

2026-09-01 developmental biology 10.64898/2026.08.31.748309 medRxiv
Top 5%
0.1%
Show abstract

Hirschsprung disease (HSCR) is a severe birth defect where ganglia of the enteric nervous system (ENS) are missing from distal bowel. The aganglionic segment is also characterized by increased epithelial permeability and pro-inflammatory immune activation. These problems may sequentially lead to translocation of gut microbes into the colon wall and systemic circulation, resulting in enterocolitis and sepsis. Current HSCR treatment via surgical resection of the aganglionic segment is lifesaving but not curative, often leaving patients with persistent gastrointestinal complications including recurrent risk of enterocolitis. As alternative, we are developing a regenerative medicine strategy based on in situ stimulation of tissue-resident ENS progenitors via rectal administration of the neurotrophic factor GDNF. Here, we report that GDNF-based therapy has pleiotropic gastrointestinal effects in a mouse model of short-segment HSCR, beyond its role in ENS regeneration. Interestingly, we found that these protective effects are not restricted to the aganglionic distal colon, also positively impacting the ENS-containing proximal colon. GDNF treatment reduces bacterial translocation both locally and in peripheral organs, and this is associated with recovery of the key epithelial junction proteins CLDN3, ZO1 and DSG2. Furthermore, multiparameter flow cytometry-based analysis of 55 lymphoid and 17 myeloid cell subtypes revealed that GDNF treatment has global anti-inflammatory effects, preferentially affecting innate over adaptive immunity. Overall, these findings highlight a critical role for GDNF treatment in reestablishing proper epithelial and immune cell homeostasis, offering promising therapeutic avenues not only for HSCR but also potentially for other intestinal disorders with overlapping pathophysiology.

14
Bayesian adaptive experimental design for efficient microbial genome-wide association studies

Helekal, D.; Blomqvist, S. O. P.; Mukherjee, A.; Bowcutt, B. A.; Palace, S. G.; Grad, Y. H.

2026-08-31 genetics 10.64898/2026.08.26.747358 medRxiv
Top 6%
0.1%
Show abstract

Bacterial genome-wide association studies (GWAS) offer a powerful approach to identify the genetic basis of a trait measured in a set of sequenced isolates. As the number of sequenced isolates has grown, the limiting factor for GWAS has become phenotyping enough isolates to achieve statistical power. To overcome the need for large-scale phenotyping, we developed Bayesian Adaptive Sequential Sampling GWAS (BASS-GWAS), which couples Bayesian adaptive experimental design with a sparse regression model to select maximally informative isolates for phenotypic testing. BASS-GWAS efficiently recovered causal loci for three antimicrobial resistance traits in Neisseria gonorrhoeae, requiring many fewer phenotyped isolates than random sampling. We applied BASS-GWAS to discover variants enabling gyrBD429N-dependent cross-resistance to the novel topoisomerase inhibitors zoliflodacin and gepotidacin. After phenotyping fewer than 30 isolates, we identified and then validated both parCD86N and a gyrA-parE-based pathway as enabling cross-resistance. BASS-GWAS provides a practical and statistically principled solution for efficient bacterial GWAS.

15
Pan-cancer Graph-based Cancer Detection Using the Cell-free DNA Methylome

Zhao, L.; Zeng, Y.; Abelman, D. D.; Lin, W.; Luo, P.

2026-08-31 oncology 10.64898/2026.08.26.26361432 medRxiv
Top 7%
0.1%
Show abstract

Motivation: Cell-free DNA methylation provides a minimally invasive signal for early cancer detection and tissue-of-origin prediction. Most methods represent methylation measurements as independent fixed-window features and therefore do not explicitly model relationships among genomic regions. Results: We developed PANGEM (Pan-cancer Graph-based Cancer Detection Using the Cell-free DNA Methylome), a graph-learning framework that represents genomic bins as nodes and integrates CpG context, genomic proximity, and sample-specific methylation similarity in the graph topology. Across five repeated stratified train-test splits, PANGEM achieved the highest mean performance among evaluated methods, with an AUROC/AUPR of 0.997/1.000 for binary cancer detection and macro-AUROC/AUPR of 0.977/0.870 for multiclass tissue-of-origin prediction. In the independent INSPIRE cohort, 72 of 78 cancer cases (92.3%) exceeded the binary classification threshold, and PANGEM correctly classified 9 of 17 head and neck cancer cases (52.9%), the highest accuracy among evaluated methods. Subnetwork analysis further identified recurrent, graph-connected methylation patterns, including a 111-DMR subnetwork with increased methylation in cancer samples.

16
Automatic bioinformatic software named entity recognition from literature

Xuan, H.; Pasupuleti, R.; Liu, B.; Sun, H.; Zhang, J.; Yao, Z.; Zhong, C.

2026-09-01 bioinformatics 10.64898/2026.08.26.731133 medRxiv
Top 7%
0.1%
Show abstract

Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale. The lack of a comprehensive and up-to-date catalog of bioinformatics resources hinders efforts toward automated biomedical knowledge extraction and streamlined data analysis. Here we present SNAIL, a hybrid named entity recognition framework designed to automatically identify bioinformatics software and database (SW/DB) names from biomedical texts. SNAIL integrates complementary lexical and semantic modeling strategies. The lexical component captures orthographic patterns and contextual cues characteristic of SW/DB names, while the semantic component leverages contextual embeddings generated by transformer-based language models such as SciBERT, combined with an explicit token-masking strategy to enhance entity-focused representations. A large training corpus was constructed automatically through a hybrid pipeline that integrates citation-hinted extraction with large language model-assisted distillation. Evaluation on two independent benchmark datasets and real-world research articles demonstrates that SNAIL substantially outperforms existing approaches, including domain-specific methods such as bioNerDS2 and general-purpose large language models such as ChatGPT, Gemini, Grok and Claude. Applying SNAIL to large-scale literature analysis further reveals distinct journal-level preferences across bioinformatics subfields. These results demonstrate that SNAIL provides an accurate and scalable solution for identifying bioinformatics resources in scientific texts and enables systematic meta-analysis of tool usage and research trends.

17
ICONIC: An R Package for Integrating Instrumental Variable- and Negative-Control-Informed Causal Discovery and Diagnostics in Multiomic Studies

Bresnahan, S. T.; Xiong, C.; Head, T.; Chang, Y.-H.; Bhattacharya, A.; Huang, J. Y.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.26.26361466 medRxiv
Top 7%
0.1%
Show abstract

Unmeasured confounding threatens causal inference and replicability in observational multi-omic studies across variable environments. Genetic instrumental variables (Mendelian randomization) and negative-control calibration each address complementary sources of unmeasured confounding, yet no existing framework unifies them for omics-scale mediation analysis. We introduce ICONIC, an R package that embeds genetic instruments and negative controls within a proximal causal inference framework for total-effect and mediation analysis. ICONIC implements eight estimators spanning five confounding-control strategies, supports continuous, binary, and time-to-event outcomes, and provides extensive diagnostics including sensitivity analyses that map estimator performance across plausible assumptions. Ground-truth benchmarks are calibrated to real-omics covariance structures via a hybrid generative model (GAN + feature-level Gaussian copula) rather than parametric simulation, and a companion planning tool predicts performance gains from collecting additional omic data. We demonstrate ICONIC in two case studies: identifying placental transcriptomic mediators of gestational diabetes on birth weight (n = 164), and tumor-expression mediators of smoking intensity on lung cancer survival (n = 494). Notably, ICONIC's diagnostics recommended different estimation strategies across the two scenarios, reflecting differences in the likely influence of unmeasured confounding. ICONIC is freely available at https://github.com/sbresnahan/iconic/.

18
A mutation-agnostic and allele-specific ASO strategy demonstrates potent functional rescue and retinal preservation in RHO-linked retinitis pigmentosa

Spaag, S.; Wu, W.-H.; Yun, J.; Winogrodzki, T.; Knudsen, A. S.; Fuso, M.; Stingl, K.; Komissarov, G.; Armento, A.; Baumann, B.; Kuehlewein, L.; Ayuso, C.; Fernandez-Caballero, L.; Collin, R.; Corradi, Z.; Roosing, S.; Kaltak, M.; Lochmann, C.; Radboudumc, F.; Banfi, S.; Karali, M.; Bolz, S.; Simonelli, F.; Dave, K.; Kohl, S.; Zrenner, E.; Demirkol, A.; Achberger, K.; Wissinger, B.; Tsang, S. H.; De Angeli, P.

2026-09-01 genetics 10.64898/2026.08.25.747013 medRxiv
Top 7%
0.1%
Show abstract

Autosomal dominant retinitis pigmentosa (adRP) caused by RHO mutations is a leading form of inherited retinal degeneration. Extensive allelic heterogeneity of RHO pathogenic variants limits the translational applicability of mutation-specific gene therapies. To address this, we developed SNARE (SNP-guided Silencing of Aberrant RHO Expression), a mutation-independent, allele-specific antisense oligonucleotide (ASO) strategy. SNARE selectively suppresses mutant RHO transcripts by targeting the common, benign c.-26A/G single-nucleotide polymorphism (SNP) as an allelic discriminator. Candidate gapmer ASOs were screened in engineered reporter lines and validated in patient-derived retinal organoids, identifying RHOligo-A as the lead c.-26A-targeting candidate. In vitro, RHOligo-A achieved robust, preferential knockdown of the target allele, improving RHO localization in retinal organoids, and demonstrated a favorable safety profile with minimal transcriptomic off-target effects and no detectable immunostimulatory activity. Subsequent validation in a novel, humanized RHOP347L/WT mouse model, achieved sustained c.-26A-linked allele-selective suppression, retinal structure preservation, and significantly restored visual function, upon a single intravitreal administration. These findings establish RHOligo-A and SNARE as a scalable, mutation-independent therapeutic platform with strong translational potential and substantial clinical reach for RHO-associated adRP.

19
Making Accelerating Medicines Partnership Data Findable and Interoperable through a Common Data Model: Extending OMOP for Multi-Source Multimodal Data

Tindall, C.; Long, R. A.; Naughton, B.; Mapes, B. M.; Vismer, D.; Skinner, H. G.; Malenfant, J.; Maurya, M. R.; Nalls, M. A.; Ramachandran, S.; Nguyen, T.; Peters, M. A.; Scheuermann, R. H.

2026-09-02 genetic and genomic medicine 10.64898/2026.08.31.26361831 medRxiv
Top 7%
0.1%
Show abstract

SysBio FAIRplex is a Common Fund Venture Program that catalogs and indexes data from the Accelerating Medicines Partnership(R) (AMP(R)) Program through a federated model in which data hosts retain custody of their datasets. The central piece of this work is the SysBio Common Data Model (SysBio CDM). AMP is a precompetitive public-private partnership started in 2014 that unites the resources of NIH and private partners to improve our understanding of disease pathways and transform current models for developing new treatments by: - identifying new targets, biomarkers, and development paradigms; - developing leading-edge tools and technologies; - collecting large-scale datasets and supporting analytics for open analysis by the public; and - generating consensus platforms and procedures. A multidisciplinary Task Force was chartered to design the SysBio CDM by extending the Observational Medical Outcomes Partnership (OMOP) Common Data Model into the -omics domain. The Task Force produced a Minimum Viable Product comprising nine OMOP tables; four extension tables for assay and file metadata; and a Common Data Element (CDE) Registry to specify field semantics. This manuscript describes the deliverable: the underlying design choices, the criteria applied in selecting and constructing the extension tables, how the extended model supports multimodal data integration across AMP projects, and what further work to support additional -omics modalities would entail. As an auxiliary methodology, the paper also describes the AI-assisted CDE harmonization workflow used to populate the model.

20
CyChat: a conversational Cytoscape app for no-code, reproducible network analysis

Liebold, J.; Stahl, M.; Schulze, J.-O.; Razavi, M. M.; Bader, G. B.; Kurtz, S.; Baumbach, J.

2026-09-01 bioinformatics 10.64898/2026.08.28.747833 medRxiv
Top 7%
0.1%
Show abstract

Network-based analyses of molecular interactions are useful for interpreting high-throughput omics data and identifying therapeutic targets. Cytoscape is the standard platform for these tasks, but users face a trade-off between accessible graphical workflows that are difficult to document and reproducible automation in Python or R that requires programming expertise. General-purpose coding assistants can generate Cytoscape Automation scripts, but remain external to Cytoscape. We present CyChat, a Cytoscape Desktop app that integrates a chat interface and a large language model (LLM) agent into the application. CyChat translates natural language into executable Cytoscape Automation workflows, runs generated Python code, and exports chat sessions with executed code as standalone Jupyter notebooks. To reduce setup barriers, CyChat includes an embedded Python runtime and supports both cloud-based and locally hosted LLMs. CyChat was evaluated across ten Cytoscape workflows using seven LLM providers, each represented by one LLM. The strongest configuration achieves a pass rate above 99%. In a qualitative evaluation based on a published network visualization, CyChat completes the task in 1.5-5 minutes, compared with 15-20 minutes for manual GUI workflows by computational biologists. CyChat is available through the Cytoscape App Store at https://apps.cytoscape.org/apps/cychat.